Ir arriba
Información del artículo

VERSE: Visual Embedding Reduction and Space Exploration - Latent-space clustering for improving document understanding

I. de Rodrigo, A.J. López López, J. Boal

Pattern Recognition Vol. 180, nº. Part D, pp. 114448

Resumen:

Conventionally, synthetic training data quality is evaluated through human perception, prioritizing visual realism. From the model’s perspective, what truly matters is whether a sample lies within the right region of its embedding space. This work introduces VERSE, a methodology for analyzing and improving the performance of Vision–Language Models by exploring their visual embedding space. VERSE enables the visualization of latent representations to assess model feasibility, identifies problematic regions, and guides synthetic data generation to enhance performance in those clusters. We validate the proposed methodology for Visually-rich Document Understanding by training on the synthetic MERIT Dataset and evaluating on its real-world counterpart, MERIT Secret, focusing on key information extraction as a sequence-generation task scoped to transcripts of records in Spanish. Results show that VERSE uncovers the visual features associated with error-prone clusters, and that retraining with samples containing these features substantially boosts F1 performance without degrading generalization. On-premise models optimized with VERSE—Donut (F1 = 0.76) and Idefics2 (F1 = 0.81)—match or surpass SaaS solutions such as GPT-4o (F1 = 0.78) and Pixtral (F1 = 0.73), preserving data privacy and avoiding external APIs.


Resumen divulgativo:

VERSE es una metodología para analizar y mejorar VLMs explorando su espacio de embeddings visuales. Identifica clústeres propensos a error y guía la generación de datos sintéticos, logrando que modelos on-premise (Donut, Idefics2) igualen o superen a soluciones SaaS como GPT-4o preservando la privacidad.


Palabras Clave: Visually-rich Document Understanding; Vision-Language Models; Visual embeddings; Interpretability; Explainability


Índice de impacto JCR-JIF y cuartil WoS: 9,100 - Q1 (2025)

Referencia DOI: DOI icon https://doi.org/10.1016/j.patcog.2026.114448

Publicado en papel: Diciembre 2026.

Publicado on-line: Julio 2026.



Cita:
I. de Rodrigo, A.J. López López, J. Boal, "VERSE: Visual Embedding Reduction and Space Exploration - Latent-space clustering for improving document understanding", Pattern Recognition, Vol. 180, nº. Part D, pp. 114448, Diciembre 2026. [Online: Julio 2026] doi: 10.1016/j.patcog.2026.114448

    Líneas de investigación:
  • Aprendizaje Profundo para la Optimización de Procesos y Activos Industriales
    Grupos de investigación:
  • Instituto de Investigación Tecnológica (IIT)
    ODS:
  • Objetivo 9: Industria, innovación e infraestructuras